Papers with speech technologies
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for learning meaningful representations from unannotated data are resource-intensive and degrade other speech components. |
| Approach: | They propose a method that decomposes SSL representations into speaker-specific components and generates speaker disentangled representations. |
| Outcome: | The proposed method achieves speaker independence and improves on state-of-the-art methods. |
Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond (2025.naacl-long)
Copied to clipboard
Mardhiyah Sanni, Tassallah Abdullahi, Devendra Deepak Kayande, Emmanuel Ayodele, Naome A Etori, Michael Samwel Mollel, Moshood O. Yekini, Chibuzor Okocha, Lukman Enegi Ismaila, Folafunmi Omofoye, Boluwatife A. Adewale, Tobi Olatunji
| Challenge: | Afrispeech-Dialog is a benchmark dataset of 50 simulated medical and non-medical African-accented English conversations . a 10%+ performance degradation is found in ASR systems on long-form, accented speech . |
| Approach: | They propose to use a dataset to evaluate automatic speech recognition systems on African-accented conversations. |
| Outcome: | The proposed dataset compares state-of-the-art speech recognition systems on accented conversations with native accents and shows a 10%+ performance degradation. |
Open-source Multi-speaker Corpora of the English Accents in the British Isles (2020.lrec-1)
Copied to clipboard
| Challenge: | Using a dataset of high-quality audio, the authors examine the accents of 120 volunteers in the British Isles. |
| Approach: | They present a dataset of high-quality audio of English sentences recorded by volunteers with different accents of the British Isles. |
| Outcome: | The transcribed audio includes pronunciations of global locations, major airlines and common personal names in different accents. |
Phonetic Segmentation of the UCLA Phonetics Lab Archive (2024.lrec-main)
Copied to clipboard
| Challenge: | ''big data'' does not exist for the majority of the world's languages . a corpus of audited phonetic transcriptions and phone-level alignments is available for free . |
| Approach: | They present a corpus of audited phonetic transcriptions and phone-level alignments from the UCLA Phonetics Lab Archive . they discuss the utility of the corpus for general research and pedagogy in crosslinguistic phonetics . |
| Outcome: | The VoxAngeles corpus improves the original corpus for phonetic typology and word- and phone duration measurements. |